JMIR Public Health and Surveillance
◐ JMIR Publications Inc.
Preprints posted in the last 30 days, ranked by how well they match JMIR Public Health and Surveillance's content profile, based on 45 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit.
Sadeghi Naieni Fard, F.; Oppong, J. R.; Tiwari, C.; Boakye, K.; Fard, F.
Show abstract
Cancer prevalence is distributed unevenly across regions and caused by the interaction of multiple risk factors. Previous studies focused on the use of global modeling techniques to predict cancer at the county level that overlooks important spatial differences. This study aims to develop geographically weighted machine learning models to predict cancer prevalence at the census tract level in the United States and identify local determinants of cancer burden. First, a scoping review was conducted to find a list of measurable drivers of cancer in the United States. Using this list, the data of these variables for 84415 census tracts were obtained from the Center for Disease Control and Prevention PLACES dataset and other publicly accessible resources. Then, several predictive models, including Ordinary Least Squares (OLS) and Geographically Weighted Regression (GWR), as well as Random Forest, XGBoost, and Deep Neural Network and their geographically weighted counterparts, were developed and compared using the Coefficient of Determination, Root Mean Square Error, and Absolute Error. Results presented that geographically weighted models outperformed other methods, and geographically weighted XGBoost achieved the strongest and most consistent overall performance with pseudo-R2 ranging between 0.89 and 0.98. Feature importance analysis of this model illustrated that most important cancer drivers changed location by location. Aged people, racial composition, preventative behaviors, and metabolic conditions such as diabetes, hypertension, and high cholesterol were determined as influential predictors, although their relative importance varied across regions. These findings revealed the value of localized models at a small geographic scale to identify regional cancer risk patterns and help the allocation of proper resources to hotspot areas. Keywords: Cancer prevalence, Census tracts, geographically weighted machine learning models, Deep neural network, XGBoost, Random Forest, Ordinary Least Squares, risk factor, determinant
Wang, K.; Olaniyan, P.; Powla, P.; Pabon-Rodriguez, F. M.
Show abstract
Indiana still faces significant health challenges, ranking among the least healthy U.S. states due to high obesity rates, mental health issues, and other chronic conditions. These disparities are closely linked to inequities in healthcare access, which are largely shaped by social determinants of health. Using data from the Social Vulnerability Index and County Health Rankings and Roadmaps, this study analyzes trends in obesity, mental health, and premature death across Indiana counties before, during, and after the COVID-19 pandemic. Descriptive statistics, correlation analyses, and Negative Binomial regression models were used to evaluate county-level disparities. In 2018, higher rates of uninsured, obese, and physically inactive populations were associated with increased premature death. In 2020, diabetes, smoking, and alcohol consumption were significant factors. By 2022, unemployment, education, obesity, insurance, exercise access, and mental health provider availability were associated with premature death. Findings indicate that socially vulnerable counties experienced amplified health impacts, with obesity rising most sharply where exercise infrastructure was limited and poor mental health days increasing across all counties. These results highlight persistent service gaps and the critical need for targeted investments in recreational infrastructure and mental healthcare. Future research should examine policy influences and causal relationships to inform equity-focused interventions.
Jaganath, D.; Ilavarasan, V.; Wong, R.; Chitnis, A.; Murrill, M. T.
Show abstract
Context: Most individuals in the United States have commercial health insurance, yet costs for tuberculosis (TB) care have focused on the public sector. Objective: To quantify 12 month all cause healthcare costs and identify predictors of expenditure among commercially insured persons with TB disease in the United States. Design/Setting: Retrospective cohort study using Merative (TM) MarketScan (R) Commercial Claims Database (2013 to 2018). Participants: Adults 18 years old with TB disease Main Outcome Measure: Total 12 month all cause healthcare costs (outpatient, inpatient, pharmacy) were calculated from the date of diagnosis. Adjusted cost ratios (aCR) were estimated using a Gamma generalized linear model. Results: We included 303 individuals diagnosed with TB disease, median age 46 years, 158 (52%) male, 16 (5%) with HIV, 12 (4%) with hepatitis B (HBV), and 13 (4%) with a drug use disorder. Mean total 12-month costs were $32,404 (median $8,075; SD $78,829). Median 12-month costs were substantially higher among persons with any comorbidity (HIV, HBV, hepatitis C (HCV), alcohol use disorder, drug use disorder, or Charlson score >0) compared to those without ($11,930 [IQR $4,194 to $36,073] vs $3,385 [IQR $1,506 to $8,609]; p<0.001). HIV coinfection and drug use disorder were the strongest independent predictors. HIV coinfection was associated with 4.7 fold higher costs (aCR 4.70, p<.001), driven predominantly by pharmacy expenditure (aCR 16.4). Drug use disorder was associated with 3.2 fold higher costs (aCR 2.62, p=.03). Comorbidity burden was a continuous independent predictor (aCR 1.36 per Charlson point, p<.001). Conclusions: Healthcare costs are high among persons with TB who have commercial insurance, and are further increased with comorbidities including HIV coinfection and drug use disorder. Improved screening, care coordination and management of TB and high risk comorbidities could yield significant cost savings.
Chaturvedi, R. R.; Gracner, T.; Perez-Arce, F.; Suen, S.-c.; Jin, J.; Orriens, B.; Pacula, R. L.; Sexton Ward, A.; Haile, R.; Kapteyn, A.
Show abstract
Importance: Evidence on GLP-1/GIP therapies is largely derived from trials enrolling selected populations or medical records that miss utilization outside healthcare channels. No nationally representative cohort has characterized real-world uptake, indications, and access. Objective: To characterize GLP-1/GIP prevalence, indication, clinical profile, and access. Design: Prospective cohort study with three GLP-1/GIP surveillance waves (March 2024, December 2024, October 2025). Setting: The Understanding America Study, an address-based, nationally representative panel of approximately 15,000 US adults aged 18+ years initiated in 2014. Participants: UAS participants responding to at least one surveillance wave (n=9150). Exposures: GLP-1/GIP use status (never vs any use, comprising current and former use), self-reported primary indication (diabetes, weight loss, or other), and access pathway (traditional vs non-traditional). Main Outcomes and Measures: Survey-weighted prevalence of GLP-1/GIP use, overall and by indication and access pathway; sociodemographic, cardiometabolic, treatment, and access characteristics; and smartwatch-derived resting heart rate, heart rate variability, maximum activity heart rate, step count, and sleep duration and variability. Results: Among n=9150 adults (1274 with any use; 60.9% female; median age 53 years), weighted prevalence increased 46%, from 8.2% (March 2024) to 12.0% (October 2025) representing 32 million. Weight-loss indications grew, reaching nearly half of use (4.1% to 5.6%); diabetes-indicated use was stable (5.3% to 5.4%). Users carried high cardiometabolic burden (obesity, 68.2%; diabetes, 53.6%) but diverged by indication: diabetes-indicated users were older (median, 59 vs 49 years), whereas weight-loss-indicated users were more often female (69.9% vs 51.3%) and healthier. One in three users (~9 million) had non-traditional access, especially in weight-loss-indicated users, of whom 33% had no conventional prescription; 41% used compounding, online, or foreign pharmacies; and, 43% lacked coverage. Non-traditional users were five times as likely to report an unlisted, likely compounded formulation (19.8% vs 4.1%). All p<0.05. Conclusions and Relevance: Real-world GLP-1/GIP use has grown rapidly and diversified substantially in indication, access, and population profile. One in 3 users obtained treatment through nontraditional channels largely invisible to claims data, raising long-term safety, efficacy, and coverage questions. GLIMMER provides a public, nationally representative longitudinal evidence base for future payer and provider decisions.
Srivastava, D. K.; Gupta, S.; Yadav, N.
Show abstract
Background: Evaluation of public health surveillance systems is a programmatic obligation but has largely been conducted as a periodic, externally commissioned activity requiring dedicated resources and additional data collection. India's Integrated Disease Surveillance Program (IDSP) generates continuous outbreak data through weekly reports but lacks a routine, embedded performance evaluation mechanism. This study assessed the quality of IDSP outbreak detection and response across multiple surveillance attributes and developed a weighted composite performance scoring framework using only routine program data. Methods: A cross-sectional evaluation study was conducted across 38 districts of Bihar using secondary data from IDSP Central Surveillance Unit weekly outbreak reports for 2016 - 2018 (n=559 outbreaks). Six surveillance quality attributes were assessed - timeliness, completeness, representativeness, relative sensitivity, acceptability and flexibility. A weighted composite performance scoring scale was developed using expert opinion-derived attribute weightages (n=25 experts). District-level scores were computed and scaled to 100. Results: Timeliness was the poorest-performing attribute, with fewer than 15% of outbreaks notified within 48 hours across all three years. Private sector participation was entirely absent - the acceptability score was 0 across all 38 districts for all three years. Completeness was the strongest attribute, exceeding 95% in all years. The mean composite score remained consistently low (23 - 27 out of 100) with widening inter-district disparity over time. Four districts (10.5%) scored 0 in all three years. Conclusions: This study presents a dynamic, routine-data-based composite performance evaluation framework for IDSP outbreak detection and response. The modular, configurable framework functions at any administrative level (from block to national) and is compatible with digital health information platforms, enabling continuous, embedded performance monitoring without additional data collection. The framework has been registered as an Intellectual Property with the Government of India. Keywords: Disease surveillance; IDSP; IDSR; performance evaluation; composite score; outbreak detection; timeliness; completeness; relative sensitivity; digital health
Davis, J. T.; Kaur, G.; Hines, A.; Ben-Nun, M.; Venkatramanan, S.; Brooks, L.; Mathis, S.; Ajelli, M.; Litvinova, M.; Kummer, A. G.; Ventura, P. C.; Mhade, S.; Weber, D.; Shemetov, D.; DeFries, N.; McDonald, D. J.; Yamana, T.; Zepeda-Tello, R.; Shaman, J.; Yaari, R.; Pei, S.; Webber, A.; Shandross, L.; Ray, E.; Wadsworth, S.; Niemi, J.; Redman, W. T.; Mullany, L.; Posner, R.; Mallela, A.; Lin, Y. T.; Hlavacek, W. S.; Smart, A.; Gill, A. A.; Drennan, A.; Fiebiger, B. J.; Miller, E. F.; Lee, J.; Mihaljevic, J. R.; Geist, K. A.; Baltz, M.; Bernik, O.; Truong, Y.-M. B.; Chen, Y.; Grosvenor, C. J.;
Show abstract
Forecasting influenza hospitalizations informs public health preparedness, yet questions remain about which types of forecasts best guide action. We evaluate categorical trend forecasts, which communicate probabilities of upcoming increases or decreases in epidemic trajectories, submitted to CDC's FluSight Forecasting Challenge between Fall-2024 and Spring-2026. Teams submitted probability distributions over five categories describing direction and magnitude of week-over-week changes in laboratory-confirmed influenza hospital admissions. We assessed performance using Ranked Probability Skill Score, Brier Skill Score, and measures of forecast-observation agreement. Most models outperformed an equal-probability baseline; the FluSight ensemble ranked among the top three in the 2024-25 and 2025-26 seasons. Forecasts were most accurate during stable periods and least during periods of rapid change, with most models underestimating observed trends. Conclusions were robust to choice of scoring metric and reference model. These results support categorical trend ensembles as an approach to communicating infectious disease forecasts that may inform public health decision-making.
Packard, S. E.; Russo, T.; Parrott, J.; Sisti, J.; Lans, A.
Show abstract
Objectives: To estimate the prevalence of Post-Exertional Malaise (PEM) among adults with prior COVID-19 and associated mental health and disability outcomes. Methods: We conducted a cross-sectional analysis of data from a survey of 9,620 adults with prior COVID-19 in New York City, collected May - June 2024. PEM was measured with the DePaul Symptom Questionnaire - Post Exertional Malaise, categorized by symptom duration (< 14 vs. [≥]14 hours). Weighted prevalence estimates were stratified by socio-demographic and clinical characteristics. Modified Poisson regression was used to assess the association of PEM with depression, anxiety, and disability. Results: The prevalence of PEM symptoms was 20.9% overall and 4.0% with symptom duration [≥]14 hours, representing over 800,000 New Yorkers affected and over 150,000 who meet a diagnostic criterion for ME/CFS. PEM prevalence was higher among women, transgender and non-binary adults, people of color, and lower educational attainment, chronic comorbidities, or disabilities. PEM was associated with 3 - 4 times higher prevalence of mental health outcomes and 4 - 5 times higher disability scores. Conclusions: PEM symptoms were common and strongly associated with disability and adverse mental health. Screening, pathways to care, and supportive policies are needed to mitigate long-term consequences, particularly among marginalized populations.
kobayashi, v.; Baluyut, G. T. C.
Show abstract
Purpose Prevention and early detection of osteoporosis remains a global challenge, more so in regions like the Philippines where screening barriers exist. Chest x-rays meanwhile are relatively inexpensive, and more frequently done, and therefore can be used for opportunistic screening. This study aimed to develop a deep learning model for osteoporosis detection from chest x-rays using DXA as the gold standard. Methods A convolutional neural network called Osteo-AI was developed using 406 pairs of chest x-rays and DXA scans of Filipino patients aged 50 and above. With data augmentation, the training set expanded to 6,300 pairs. Gradient-weighted class activation mapping technique was applied to localize and identify patterns and areas in the chest x-ray images correlating with osteoporosis. Results Training data consisted of 369 female patients and 37 males. Ages of the patients ranged from 50 to 89 with a mean age of 63 years old. Initial testing yielded promising results, with Osteo-AI achieving a diagnostic accuracy of 85.71%, easily outperforming a benchmark of 33.33% Conclusion Our findings suggest the potential of Osteo-AI to enhance osteoporosis screening accessibility, aiding in early intervention to prevent fragility fractures. Further research involving larger datasets is warranted to refine and optimize the model, potentially improving detection accuracy and expanding its utility in global healthcare settings.
Amolo, P.; Mungai, L.; Karume, A. K.; Kibugi, J.; Mwende, W.; Botella, N.; Haldane, C.; Kamau, Y.; Marban-Castro, E.
Show abstract
Introduction Continuous Glucose Monitoring (CGM) is considered standard care in high-income countries. There is, however, limited published evidence on CGM use in low- and middle-income countries. The purpose of this study was to assess the usability, acceptability, and feasibility of CGM use among people living with type 1 diabetes (T1D) and caregivers in a low-resource setting. Research Design and Methods This prospective study conducted at the Kenyatta National Hospital purposively enrolled persons aged 4-25 years who had been on management for T1D for at least six months, and caregivers of those under 18 years. Fourty youth living with T1D used CGM for three months in place of self monitoring of blood glucose (SMBG). The System Usability Scale (SUS), a Theoretical Framework of Acceptability-based questionnaire, the Diabetes Distress Scale (DDS), the Glucose Monitoring Satisfaction Survey (GMSS), and a feasibility survey were administered. Outcomes were summarized descriptively, including means, medians, and frequencies using R statistical software. Results The median SUS score was 98.8 (IQR 92.5-100.0). Acceptability was high, and the median total GMSS score improved from 3.73 to 4.73. Among adolescents and adults, the median overall DDS score reduced from 1.54 to 1.36, with reductions in scores in all domains, except for hypoglycemia distress which increased, and physician distress which remained low. Among caregivers, the median overall DDS score declined from 2.05 (moderate distress) to 1.90 (low distress), with modest reductions in teen management and parent-teen relationship distress and a slight increase in personal distress. Median CGM active wear time was 89%. Conclusion This study comprehensively evaluated CGM across usability, acceptability, and feasibility outcomes, with the findings supporting the integration of CGM into routine diabetes management in low-resource settings. The short follow-up period, however, may not capture changing perceptions or long-term adherence.
Karume, A. K.; Amolo, P.; Mungai, L.; Moraa, H.; Arunga, T.; Nzove, E.; Ndambuki, C.; Muhwava, L.; Kamau, Y.; Marban-Castro, E.
Show abstract
Background: Type 1 diabetes (T1D) is a growing public health concern in low- and middle-income countries, where access to glucose monitoring and consistent routine care remains limited. Continuous glucose monitoring (CGM) may improve diabetes outcomes, but evidence of its acceptability and use in low-resource settings is limited. This study explored the perceptions and experiences of CGM use among people living with T1D, their caregivers, and healthcare providers (HCPs). Methods: Participants were recruited from the ACCEDE-U study, a usability study on CGM use among people living with T1D attending a tertiary referral hospital in Nairobi, Kenya. Among 40 participants in the ACCEDE-U study, those who had completed at least seven weeks of CGM use were eligible to participate in the qualitative component. Three focus group discussions (FGDs) were conducted: one with nine caregivers, one with nine adolescents (12-17 years), and one with six young adults (18-24 years). Semi-structured interviews were conducted with 9 HCPs. Data was collected using guides, audio-recorded, transcribed, and analyzed thematically guided by the socioecological framework. CORE-Q guidelines were used to report results. Results: Participants reported increased engagement in glucose monitoring and high acceptability of CGM. Reduced finger-prick testing and real-time alerts were key benefits, particularly among adolescents and young adults, who valued its discreetness and convenience. CGM was perceived to facilitate sharing of glucose data with HCPs. Caregivers reported a reduced monitoring burden. HCPs perceived CGM as valuable for clinical decision-making by providing real-time insights into glycemic patterns. Cost and limited device availability were identified as major barriers to sustained use. Conclusion: CGM was well accepted and perceived as beneficial. However, challenges related to cost and access may limit broader uptake. Improving affordability and availability could enhance feasibility and promote wider implementation in similar settings.
Maleki, C.; Bertrand, Y.; Gailly, F.
Show abstract
Clinical recommendations are often expressed in narrative form, which limits their direct execution, auditability, and patient-specific interpretation. This paper presents a hybrid decision-support framework that combines Decision Model and Notation (DMN), survey-weighted rule-ensemble learning, and counterfactual sensitivity analysis. The framework is evaluated using an NHANES-derived fasting cohort for classification of documented diabetes status. The full fasting analysis cohort contained 2,582 participants, and a non-diagnostic laboratory subgroup, Gate0, contained 2,111 participants. On untouched test data, the rule-ensemble model achieved ROC-AUC and PR-AUC values of 0.959 and 0.873 in the full fasting cohort and 0.861 and 0.499 in Gate0. Four clinically interpretable candidate rules were selected using validation data only. A nonnegative survey-weighted logistic model removed one redundant rule and converted the remaining three binary activations into an auditable DMN score and model-estimated probability. The final DMN achieved ROC-AUC 0.769, PR-AUC 0.153, and Brier score 0.029 in the untouched Gate0 test set. In small rule-defined test subgroups, hypothetical five-unit BMI reductions lowered mean model-estimated probability by 2.40 to 5.89 percentage points when one or more BMI thresholds were crossed. These findings characterize policy sensitivity rather than causal effects and require external validation.
Mao, Y.; Lin, J.; Zhou, A.; Zeng, S.; Yang, D.; Lin, W.; Wen, J.; Yang, W.; Chen, G.
Show abstract
Background Existing insulin resistance (IR) indices are predominantly developed in diabetic cohorts, limiting their generalizability. We developed a novel deep neural network-derived IR index (DNN-IR) using a Mixture-of-Experts (MoE) framework and evaluated its predictive performance for incident cardiovascular disease (CVD) and mortality in general populations. Methods We utilized data from three cohorts: the cross-sectional REACTION study (Fujian subcohort, 2011-2012) for DNN-IR derivation and internal validation; and two prospective cohorts, NHANES (1999-2018, linked to the National Death Index) and CHARLS (2011-2018), for external validation. The DNN-IR was developed using a deep learning model based on a Mixture-of-Experts (MoE) architecture, trained on the REACTION dataset. We evaluated the DNN-IR's utility in predicting incident CVD, cardiovascular mortality, and non-cardiovascular mortality among 13,889 NHANES and 7,047 CHARLS participants. Predictive performance was assessed via the area under the receiver operating characteristic curve (AUC). Multivariable logistic regression, restricted cubic splines, and Kaplan-Meier analyses characterized the associations between DNN-IR and clinical outcomes. Results In the REACTION cohort, DNN-IR demonstrated superior predictive performance for atherosclerotic outcomes, achieving AUROCs of 0.89 (training) and 0.84 (internal validation). In the external CHARLS cohort (median follow-up: 7 years; 1,135 incident CVD cases [16.1%]), DNN-IR yielded AUROCs of 0.72 for incident CVD and 0.77 for all-cause mortality. Fully adjusted models showed that each 1-SD increment in DNN-IR was associated with a 23% higher CVD risk (OR=1.23, 95% CI: 1.14-1.32), exhibiting a predominantly linear dose-response relationship (P-nonlinearity=0.453). In NHANES, DNN-IR robustly predicted cardiovascular (AUROC=0.77) and all-cause mortality (AUROC=0.72), alongside specific mortalities like diabetes (0.91), Alzheimer's disease (0.88), and kidney disease (0.96). Higher DNN-IR levels correlated with stepwise increases in cumulative mortality (log-rank P<0.001). Conclusions The MoE-derived DNN-IR index demonstrated robust and stable performance in predicting atherosclerosis, incident CVD, cardiovascular mortality, and all-cause mortality in the general population. Further validation in larger, more diverse cohorts is warranted to support its broad clinical applicability.
Chugh, M.; Neekhra, B.; Bamrotiya, M.; Clipman, S. J.; Gupta, D.
Show abstract
Antiretroviral therapy (ART) stock-outs interrupt treatment, increase the risk of virologic failure and drug resistance, and erode the population-level benefits of viral suppression. India's National AIDS Control Organization (NACO) manages one of the world's largest public ART programmes, where regimen transitions, evolving formulations, changing treatment guidelines, and procurement-driven fluctuations in drug consumption complicate forecasting. We developed an end-to-end, regimen-specific forecasting workflow to support procurement planning during such periods of instability. We analyzed monthly national ART consumption data from January 2013 through December 2024. A privacy-preserving synthetic dataset was used for pipeline development, followed by final evaluation on real national consumption time series. We compared three model classes, comprising five models: (1) classical models (Holt-Winters and ARIMA), (2) transformer models (TimesFM, which is a large pre-trained time-series foundation model, and its variant with logarithmically transformed values), and (3) hybrid models (variants of a hybrid ARIMA-TimesFM residual model). While the forecast horizon of 18 months remained constant, the train-test period varied across real and synthetic data, as real data was only available until February 2024. For synthetic data, models were trained through June 2023 (test window was July 2023-December 2024), while for real data, models were trained through August 2022 (our test window was September 2022-February 2024). We reported signed percentage deviation to preserve whether models tended to over-or under-predict, and selected models by the smallest absolute deviation. We then derived a regimen-specific model-error buffer, applied only to held-out under-prediction, and deployed the workflow through a no-code dashboard. Forecasting performance was determined using signed percentage deviation (SPD), wherein positive change represents under-prediction and negative change represents over-prediction. Performance varied across regimens, indicating that no single approach was best-performing for all formulations. On synthetic benchmark data, the smallest absolute deviations ranged from 0.46% for adult ABC+3TC to 11.92% for adult AZT+3TC. On real consumption data, classical methods remained competitive for some series, whereas transformer and hybrid models produced better predictive outcomes for others. For instance, for adult AZT+3TC, the Hybrid 70th percentile achieved an SPD of -2.02%, in contrast to the error range of [-15.7, 8.87] for other models. For adult Ritonavir, the ARIMA-TimesFM hybrid at the 30th percentile achieved an SPD of -5.2%, in contrast to the error range of [-14.94, 17.25] for other models. Several formulations, particularly low-volume and transition regimens, nevertheless remained difficult to forecast accurately, underscoring persistent operational uncertainty. This was especially evident across the three pediatric regimens, where all models deviated systematically in the same direction - a more concerning pattern than mere magnitude. For pediatric ABC+3TC, all models over-predicted within a narrow band of [-82.74, -67.43], while for pediatric AZT+3TC and LPV/r 125 mg, all models under-predicted, with ranges of [24.93, 63.73] and [18.32, 52.07] respectively. These findings support a portfolio approach to forecasting in national HIV programmes. Rather than replacing established public-health procurement systems, regimen-specific model selection, directional error reporting, and cautious model-error buffering can strengthen decision support during regimen transitions and other periods of unstable demand.
Ulm, C.; Golden, S. D.; Hill, F.; Wiesen, C. A.; Mills, S. D.
Show abstract
Introduction Smoking prevalence remains higher in rural than in urban populations in the United States. To examine recent trends, we assessed state-level differences in cigarette smoking between urban and rural areas from 2018 to 2024. Methods Using repeated cross-sectional data from the Behavioral Risk Factor Surveillance System, we estimated state-specific logistic regression models to examine the relationship between urban-rural county residence and cigarette smoking. Unadjusted models (model 1) included urban-rural county status and year. Subsequent models (model 2) added age, sex, and race/ethnicity. A final model (model 3) included education and an interaction term between urban-rural county status and year to examine whether gaps in urban-rural smoking changed over time. In states with significant interactions, simple effects tests compared trends for urban-rural groups separately. Results Compared to urban adults, rural adults had higher unadjusted odds of cigarette smoking (odds ratio [OR] range:1.07-1.88) in 88.4% (38/43) of states. Adjusting for demographic covariates (model 2) increased the proportion of states with significant marginal effects of rurality to 90.7% (ORs:1.09-1.87). A final model that also controlled for education (model 3) decreased the proportion of states with significant marginal effects of rurality to 60.5% (ORs:1.10-1.54). Among the 14 states with significant interaction terms, the odds of smoking declined faster among urban than rural residents. Conclusion Urban-rural differences in smoking persist across most states. No state showed a reduction in urban-rural disparities over time, and the urban-rural gap widened in 14 states. Demographic variation accounted for some, but not the majority, of observed urban-rural differences.
Edmond, E. C.; Dreyer, A. J.; Winston, A.; Khoo, S. H.; Joska, J.; Nightingale, S.
Show abstract
Background Computerised cognitive testing may address the global challenge in identifying cognitive changes in people living with HIV scalably and affordably. We assessed a computerised battery (CB) of cognitive tests, in a prospective cohort (CONNECT) of people with HIV in a low-income peri-urban area of Cape Town, South Africa during a national programmatic switch from efavirenz- to dolutegravir-based antiretroviral therapy (ART). Methods We recruited 170 people with HIV and 91 people without HIV (controls) (140[82%] and 41[45%] followed up). The CB and gold-standard pen&paper cognitive testing (P&P) were performed at both timepoints. Technology familiarity/use questionnaire data were also collected. We compared performance in detecting lower group-level cognitive performance associated with efavirenz treatment. Furthermore, the CB was compared to P&P in classifying individuals with low cognitive performance, correlation of global test scores and domain-level scores between batteries, and practice effects between timepoints. Exploratory principal component analysis was also performed. Results People with HIV on efavirenz at baseline had lower performance on the computerised battery than controls, {Delta}T=2.6, p=0.0047. This difference was lost after switching to dolutegravir-based ART at follow-up. CB and P&P global T were moderately correlated (R2=0.203, p<0.001), and the CB performed moderately in classification of low cognitive performance against the gold standard (AUC 0.70, sensitivity 0.52, specificity 0.76, PPV 0.40, and NPV 0.84). Selecting the first three principal components improved both classification of low cognitive performance (AUC 0.77) and correlation strength with P&P global T (R2=0.3, p<0.001). The CB did not show practice effects. Most participants owned a mobile phone (95%, 85.9% of these smartphones). Performance was better in smartphone owners ({Delta}T=1.8) and computer owners (23%, {Delta}T=1.8). Conclusions Delivering computerised cognitive testing was feasible in this low-income southern African setting. The CB showed reasonable construct validity (detecting known lower cognitive performance associated with efavirenz-ART) and may detect broad cognitive characteristics such as processing speed and accuracy. However, correlation of CB results with gold standard P&P testing was low-moderate and may limit its applicability as a diagnostic tool. This might be improved by including a wider range of cognitive domains tested in the CB, or data driven analysis. Brief CBs may fulfil an initial screening role, followed by more detailed clinical assessment.
Kallis, K.; Quitevis, C. R.; Ramsis, M.; Kabutey, N.-K.; Conte, M. S.; Rowe, V. L.; Humphries, M. D.; Hernandez-Boussard, T.; B. Malas, M.; Ross, E. G.
Show abstract
Background Peripheral artery disease (PAD) is a major cause of cardiovascular events but remains underdiagnosed. Electronic health record (EHR)-based machine learning models show promise for earlier detection, but developing generalizable and fair models across diverse populations remains challenging. Methods Using the University of California Health Data Warehouse, containing EHR data from five health systems, we identified patients with and without PAD. We used unsupervised clustering to define PAD phenotypes and trained a LightGBM classifier using 14,023 features spanning demographics, comorbidities, medications, laboratory values, healthcare utilization, and diagnosis, procedure, and medication codes. We evaluated performance overall and across demographic groups and phenotypes, and assessed fairness using selection rates and subgroup differences in true- and false-positive rates. Results The study included 33,739 cases and 33,739 matched controls. Clustering identified four phenotypes: patients with limited healthcare documentation (cluster 1), younger patients with severe metabolic disease (cluster 2), patients with a traditional atherosclerotic risk profile (cluster 3), and frail elderly patients with multimorbidity (cluster 4). Overall, the model demonstrated consistent performance across institutions (AUROC 0.76?0.79; AUC-PR 0.76?0.79) with well-calibrated probabilities. Performance was similar across genders, with modest variation by race and age, and was stronger in clusters 2?4. Cluster 2 demonstrated the highest sensitivity (TPR 0.87, 95% CI 0.87?0.88), while cluster 1 showed the lowest performance (TPR 0.40, 95% CI 0.39?0.41). Conclusions The EHR-based PAD detection model demonstrated consistent performance across five health systems. Phenotypic clustering revealed clinically meaningful differences in model performance adding an additional consideration in ML fairness and performance evaluations.
McLean, K. W.; LaBonte, J.; Macaulay, K.; Kassam-Adams, S.
Show abstract
This study documents the derivation and validation of a deterministic algorithm for cause-of-death (COD) ascertainment from longitudinal real-world medical claims data, evaluated against an independent state-level death certificate file. Death certificates are the dominant reference standard in mortality research but carry well-documented limitations, including primary-cause error rates estimated at 20-40\% across empirical studies. A matched analytic cohort of 216,382 individuals (Connecticut death records, 2017--2025, age 25 and above) was constructed after exclusion of mechanism-of-injury cases and removal of ill-defined symptom-code entries from both sources. Concordance between algorithmic and certificate-based COD was assessed through three complementary frameworks: age-stratified positive predictive value (PPV) at the ICD-10-CM chapter level under a full-set concordance scenario; mean absolute rank difference (MARD) for chapters identified by both sources; and analyses of breadth, depth, and code-level specificity of COD reporting. Chapter-level PPV was strongest for individuals aged 55 and above, with all estimates representing conservative lower bounds given the known error rate of the certificate reference standard. The algorithm consistently reported broader and more granular contributing cause profiles than the death certificate, with discordances directionally consistent with the well-documented tendency of certificates to under-report contributing conditions. These findings support the conclusion that algorithmic COD ascertainment from longitudinal claims data is a feasible and scalable alternative to certificate-based attribution and, at population scale, a principled methodology for characterising death certificate error rates beyond what small-sample chart review studies can achieve.
Pavia, M. J.; Amaro, I. F.; Xu, D.; Gonzalez-Hernandez, G.; Scotch, M.
Show abstract
Influenza vaccine effectiveness (VE) is estimated from a limited number of clinics using a test-negative design. These standard estimates face geographic, temporal, and operational constraints. Using Twitter/X data, we applied few-shot chain-of-thought prompting to identify self-reported vaccination status and influenza test results, then implemented a test-negative-like design to estimate VE. Our estimates fell within the range of interim reports and could complement current systems, improving feasibility, timeliness, and scalability.
Ahmed, T.; Asif, M. R. A.
Show abstract
Depression has become a serious concern for students worldwide. Aligned with the WHO Helping Adolescents Thrive (HAT) Guidelines and the Social Determinants of Health (SDoH) model, this study isolates five academically relevant factors academic pressure, work/study hours, study satisfaction, sleep duration, and financial stress from a dataset of 27,880 university students in India and quantifies their associations with depression. Unlike prior work that maximises classification accuracy, this study prioritises interpretability: logistic regression provides odds ratios (OR) with 95% confidence intervals, Random Forest (RF) and XGBoost rank predictors by feature importance, and SHAP (SHapley Additive exPlanations) values extend the analysis to individual-level risk explanation. SMOTE oversampling was applied exclusively to the training set, and performance was evaluated on the original imbalanced test set (n = 5,576). Both ensemble models achieve approximately 77-78% accuracy and an AUC of 0.845, confirmed by 5-fold pipeline cross-validation (CV AUC [~] 0.843). Academic pressure is the dominant risk factor (OR = 2.271; RF importance = 0.481; mean |SHAP| = 0.174), while study satisfaction (OR = 0.796) and sleep duration (OR = 0.835) are protective. The RF model yields a tipping point at academic pressure > 4.02, and interaction plots reveal how depression risk is amplified by low sleep, high financial stress, and extended study hours. These findings provide data-driven thresholds aligned with WHO-endorsed modifiable determinants to support early detection and institutional counselling.
Yehoshua, A.; Lupton, L. L.; Hu, T.; Cappelleri, J. C.; Gavaghan, M. B.; Puzniak, L.; Brathwaite, R.; Di Fusco, M.; Sun, X.
Show abstract
Background To characterize Coronavirus disease 2019 (COVID-19) symptom severity, and recovery from pre-infection through one month, overall and by risk groups. Methods Symptomatic adults aged [≥]18 years with test-confirmed COVID-19 were enrolled from ambulatory care clinics within a national U.S. retail pharmacy network between 10/24/2024 and 08/29/2025 (NCT05160636). Adjusted mixed models for repeated measures estimated least-squares mean changes (LSE) and standard errors (SE) from pre-infection and on Days 1-7, 10, 14, and Week 4 from enrollment in composite symptom scores (sum of severity ratings (0-3) across 14 symptoms), counts of mild-to-severe, moderate-to-severe, and severe symptoms, overall and by age and clinical risk status. Effect sizes (ES) were defined as small (0.2-<0.5), medium ([≥]0.5), and large ([≥]0.8). Results The analysis included 608 adults. On Day 1, symptom severity rose sharply from pre-infection for the composite symptom score (LSE 14.2 [SE 0.3]; ES 2.22), mild-to-severe (7.6 [0.1]; 2.72), moderate-to-severe (5.0 [0.2]; 1.77); and severe (1.8 [0.1]; 0.92) (all p<0.001). By Week 4, composite score (0.7 [0.2]; 0.26), mild-to-severe (0.5 [0.1]; 0.23); moderate-to-severe symptoms (0.1 [0.1]; 0.17) and severe symptoms (0.2 [0.1]; 0.5) remained slightly above baseline (all p[≤]0.025). Elevated severe symptom durations varied: high-risk adults (through Day 3), adults <50 years (through Day 7), and adults [≥]50 years (through Day 7). Conclusions COVID-19 was associated with notable acute symptoms in outpatients, followed by gradual improvement over time, although symptoms still persisted at four weeks. Improvement in severe symptoms varied by individual risk profile, reinforcing the importance risk-based follow-up and ongoing monitoring.